Papers with BDI dataset
Lost in Evaluation: Misleading Benchmarks for Bilingual Dictionary Induction (D19-1)
Copied to clipboard
| Challenge: | a quarter of the data consists of proper nouns, which can be hardly indicative of BDI performance, and there are pervasive gaps in the gold-standard targets. |
| Approach: | They examine the composition and quality of test sets for five different languages . they suggest future research avoids drawing conclusions from quantitative results . |
| Outcome: | The results show that a quarter of the data consists of proper nouns, which can be hardly indicative of BDI performance, and there are gaps in the gold-standard targets. |